Duplicate Removal for Candidate Answer Sentences

نویسندگان

  • Yuan Shen
  • Gabriel Zaccak
  • Boris Katz
  • Yuan Luo
  • Ozlem Uzuner
چکیده

In this paper, we describe the duplicate removal component of Infolab’s1 question answering system that contributed to CSAIL’s entry of TREC-152 Question Answering track. The goal of the Question Answering Track is to provide short, succinct answers to English sentences posed by users. In answering definition questions, we are asked to retrieve new and relevant information, in the form of short sentences or fragments from newswire text. Because many news articles overlap in content, we need to employ a duplicate removal method before presenting the results to the users. Here we present two different approaches to duplicate removal. Our first approach uses the BLEU score, a commonly used metric for machine translation evaluation, as the similarity metric between sentences. Our second approach takes a list of candidate answers, and clusters answers using word-level edit-distance as the similarity metric; the best answer from each cluster is chosen as the representative. In this paper we compare these two approaches and determine their relative performances in the duplicate detection task.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Answerfinder: Question Answering by Combining Lexical, Syntactic and Semantic Information

We present a question answering system that combines information at the lexical, syntactic, and semantic levels, in the process to find and rank the candidate answer sentences. The candidate exact answers are extracted from the candidate answer sentences by means of a combination of information-extraction techniques (named entity recognition) and patterns based on logical forms. The system part...

متن کامل

Improving Question Answering Sentence Selection by Rank Propagation

Open-domain Question Answering (QA) systems typically leverage an answer selection component to rank candidate answer sentences based on how likely they will contain the answer to a given question. This component plays a crucial rule in the QA system as it usually dictates how downstream processing modules (e.g., answer extraction) retain and present answers to users. Most existing works in thi...

متن کامل

Question Answering by Searching Large Corpora With Linguistic Methods

In this paper we describe the QuALiM Question Answering system which uses linguistic analysis of questions as well as candidate sentences in its answer finding process. To this end we have developed a rephrasing algorithm based on linguistic patterns that describe the structure of questions and candidate sentences and where precisely to find the answer in the candidate sentences. With this meth...

متن کامل

A Semantic Approach to Extract the Final Answer in SBUQA Question Answering System

In this paper we introduce the semantic approach of the answer extraction component of a question answering system called SBUQA. The answer extraction component gets the retrieved documents from a search engine according to a query (here NL question) as input. Then it represents both the question and candidate sentences using Lexical Functional Grammar (LFG), a meaning based grammar that analys...

متن کامل

ILQUA at TREC 2006

This year, we made changes to the passage/sentence retrieval component of ILQUA in handling factoid and list questions. All the other components remain same. ILQUA is an IE-driven QA system. To answer “Factoid” and “List” questions, we apply our answer extraction methods on NE-tagged passages or sentences. The answer extraction methods adopted here are surface text pattern matching, n-gram prox...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2006